Papers with multimodal language model

5 papers
The Impact of Auxiliary Patient Data on Automated Chest X-Ray Report Generation and How to Incorporate It (2025.acl-long)

Copied to clipboard

Challenge: Traditionally, CXR report generation relies on data from a patient’s exam, overlooking valuable information from patient electronic health records.
Approach: They propose to integrate patient data from ED records into multimodal language models that embed patient data into a language model.
Outcome: The proposed model incorporates patient data from the MIMIC-CXR and MIMICIV-ED datasets to improve diagnostic accuracy and improves radiologist effectiveness.
Thesis Proposal: Detecting Empathy Using Multimodal Language Model (2024.eacl-srw)

Copied to clipboard

Challenge: Existing studies on empathy detection in video and audio have relied on scripted or semi-scripted interactions that fail to capture the complexities and nuances of real-life interactions.
Approach: They propose to develop a multimodal language model that detects empathy in audiovisual data by using neural architecture search and optimisation techniques.
Outcome: The proposed model will be able to detect empathy in audiovisual data and use neural architecture search to deliver it.
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues (2025.acl-long)

Copied to clipboard

Challenge: Existing large language models fail to incorporate nonverbal elements into conversational experiences.
Approach: They propose a multimodal language model that generates nonverbal cues alongside text . their dataset is annotated with time-aligned text, facial expressions, and body language .
Outcome: The proposed model generates nonverbal languages and text, corresponding to conversational input.
Exploring Logographic Image for Chinese Aspect-based Sentiment Classification (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for aspect-based sentiment classification have focused on English text, but Chinese is a language derived from pictographs and different from other phonetic languages.
Approach: They propose to use a logographic image to capture internal morphological structure from character sequence . they propose to explicitly incorporate a symbolic image with review text for sentiment classification .
Outcome: The proposed method improves over baselines and improves on existing methods.
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling (2024.acl-long)

Copied to clipboard

Challenge: Existing language models that use discrete representations for unified processing of various modalities are limited to text generation and do not include multimodal output.
Approach: They propose a multimodal language model that utilizes discrete representations for unified processing of various modalities.
Outcome: The proposed model can be trained stably without any alterations to existing models or training paradigms.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations